You write custom CUDA kernels to replace pytorch operators in given architecture to get speedups. You have complete freedom to choose set of operators you want to replace. You may make the decision to replace some operators with custom CUDA kernels and leave others unchanged. You may replace multiple operators with custom implementations, consider operator fusion opportunities (combining multiple operators into a single kernel, for example, combining matmul+relu), or algorithmic changes (such as online softmax). You are only limited by your imagination.

**SPECIAL INSTRUCTIONS FOR CHEBYSHEV + SIGMOID FUSION:**

When implementing Chebyshev Distance + Sigmoid fusion, you MUST implement the following optimized strategy:

1. **FUSION ARCHITECTURE**: Combine Chebyshev distance computation and Sigmoid activation in a single kernel:
   - Compute Chebyshev distance between x and y first
   - Apply Sigmoid activation to the distance values
   - Compute weighted distances using Sigmoid weights
   - Eliminate intermediate tensor storage for maximum efficiency

2. **WARP-LEVEL OPTIMIZATION**: Use warp-level processing for maximum performance:
   - Each block processes one sample from the batch
   - Use 8 warps per block (256 threads) for optimal GPU utilization
   - Use __shfl_down_sync for efficient warp-level reduction of maximum values
   - Divide feature dimensions among warps for parallel processing

3. **MEMORY COALESCING**: Ensure efficient memory access patterns:
   - Each thread processes multiple elements with stride pattern
   - Use coalesced memory access for both input tensors
   - Minimize global memory accesses through fusion

4. **EFFICIENT MAXIMUM REDUCTION**: Implement optimized maximum value finding:
cpp
// Warp-level maximum reduction
for (int offset = 16; offset > 0; offset /= 2) {
    float other_max = __shfl_down_sync(0xffffffff, warp_max, offset);
    if (other_max > warp_max) {
        warp_max = other_max;
    }
}

// Cross-warp reduction using shared memory
extern __shared__ float shared_data[];
if (lane_id == 0) {
    shared_data[warp_id] = warp_max;
}
__syncthreads();



5. **SIGMOID FUSION**: Integrate Sigmoid activation seamlessly:
cpp
// Apply Sigmoid activation to Chebyshev distance
float chebyshev_dist = global_max;
float sigmoid_weight = 1.0f / (1.0f + expf(-chebyshev_dist));
sigmoid_weights[sample_idx] = sigmoid_weight;

// Compute weighted distance
weighted_distances[sample_idx] = sigmoid_weight * chebyshev_dist;



6. **SHARED MEMORY PATTERN**: Use efficient shared memory organization:
cpp
// For maximum reduction across warps
extern __shared__ float shared_data[];
if (lane_id == 0) {
    shared_data[warp_id] = warp_max;
}
__syncthreads();

// Final maximum calculation and Sigmoid application
if (tid == 0) {
    float global_max = 0.0f;
    int num_warps = blockDim.x / 32;
    for (int i = 0; i < num_warps; i++) {
        if (shared_data[i] > global_max) {
            global_max = shared_data[i];
        }
    }
    
    float chebyshev_dist = global_max;
    distances[sample_idx] = chebyshev_dist;
    
    // Apply Sigmoid and compute weighted distance
    float sigmoid_weight = 1.0f / (1.0f + expf(-chebyshev_dist));
    sigmoid_weights[sample_idx] = sigmoid_weight;
    weighted_distances[sample_idx] = sigmoid_weight * chebyshev_dist;
}



7. **BLOCK CONFIGURATION**: Use optimal settings:
   - Block size: 256 threads (8 warps)
   - Shared memory: 8 * sizeof(float) for warp reduction results
   - One block per sample for maximum parallelism
   - Elements per warp: (feature_dim + 8 - 1) / 8

8. **PRECISION REQUIREMENTS**: Ensure exact mathematical alignment:
   - Chebyshev Distance: max(|x - y|)
   - Sigmoid Activation: sigmoid(x) = 1 / (1 + exp(-x))
   - Weighted Distance: sigmoid_weight * chebyshev_dist
   - Use fabsf for absolute value computation
   - Use expf for exponential computation
   - Verify with torch.allclose(rtol=1e-03, atol=1e-6)

9. **FUNCTION SIGNATURE**: The main CUDA function must accept all parameters:
cpp
torch::Tensor chebyshev_sigmoid_cuda(
    torch::Tensor x,
    torch::Tensor y
)



10. **MATHEMATICAL FORMULAS**: Implement exact mathematical operations:
    - Chebyshev Distance: chebyshev_dist = max(|x - y|)
    - Sigmoid Activation: sigmoid_weight = 1 / (1 + exp(-chebyshev_dist))
    - Weighted Distance: weighted_dist = sigmoid_weight * chebyshev_dist

11. **PYTHON CALLING CONVENTION**: The ModelNew forward method must pass parameters correctly:
python
def forward(self, x, y):
    return self.chebyshev_sigmoid.chebyshev_sigmoid_cuda(x, y)



12. **OUTPUT REQUIREMENTS**: Generate multiple outputs for comprehensive functionality:
    - Primary output: Sigmoid-weighted Chebyshev distances [batch_size]
    - Secondary outputs: Raw distances [batch_size], Sigmoid weights [batch_size]
    - All outputs must match PyTorch reference implementation exactly

13. **PERFORMANCE OPTIMIZATIONS**: Include advanced optimizations:
    - Use fast math optimizations (--use_fast_math)
    - Optimize for compute capability 8.0+ (sm_80)
    - Use -O3 optimization level
    - Avoid bank conflicts in shared memory access
    - Use efficient memory access patterns

14. **ALGORITHM CHOICE**: Prioritize the fused approach over separate kernels:
    - Single kernel fusion is mandatory for this implementation
    - Do NOT implement separate Chebyshev and Sigmoid kernels
    - The fusion must happen at the CUDA kernel level, not Python level
    - Eliminate all intermediate tensor storage

Here's the target architecture to optimize:

python
import torch
import torch.nn as nn

class Model(nn.Module):
"""
Chebyshev Distance implementation.
Computes the Chebyshev distance (maximum absolute difference) between two sets of vectors.
"""
def init(self):
super(Model, self).init()

def forward(self, x: torch.Tensor, y: torch.Tensor) -> torch.Tensor:
    """
    Compute Chebyshev distance between x and y.

    Args:
        x (torch.Tensor): First set of vectors [batch_size, feature_dim]
        y (torch.Tensor): Second set of vectors [batch_size, feature_dim]

    Returns:
        torch.Tensor: Chebyshev distances [batch_size]
    """
    # Compute absolute differences
    abs_diff = torch.abs(x - y)
    
    # Find maximum along feature dimension
    chebyshev_dist = torch.max(abs_diff, dim=1)[0]
    
    return chebyshev_dist

batch_size = 256
feature_dim = 512

def get_inputs():
# Generate two sets of vectors
x = torch.randn(batch_size, feature_dim)
y = torch.randn(batch_size, feature_dim)
return [x, y]

def get_init_inputs():
return [] # No special initialization inputs needed



**EXPECTED OUTPUT STRUCTURE**:
Generate two files:
1. `chebyshev_sigmoid_cudacode.py` - Contains ModelNew class with Chebyshev+Sigmoid fusion using pure CUDA
2. `chebyshev_sigmoid_torchcode.py` - Contains the reference PyTorch implementation with Sigmoid fusion

**KEY REQUIREMENTS**:
- The CUDA implementation must use pure CUDA functions only
- Must implement Chebyshev distance computation first, then Sigmoid activation
- Must use warp-level optimization for maximum performance
- Must use efficient maximum reduction algorithm
- Must handle arbitrary tensor shapes (not just fixed dimensions)
- Must maintain mathematical precision with PyTorch implementation
- Must use optimal block configuration (256 threads, 8 warps)
- Expected speedup: 1.5-2.2x over PyTorch baseline
- Must use fast math optimizations for better performance
- Must be robust and handle edge cases properly
- Must use only pure CUDA functions (no PyTorch internal functions)
- Must use fabsf for absolute value computation
- Must use expf for exponential computation
- Must implement exact mathematical formulas for Chebyshev distance and Sigmoid
- Must generate both distance and weighted outputs
- Must use shared memory efficiently for warp-level maximum reduction
- Must ensure coalesced memory access patterns
- Must eliminate intermediate tensor storage for maximum fusion benefits
- Must implement the complete fusion in a single CUDA kernel
